Serveur d'exploration sur l'OCR

Attention, ce site est en cours de développement !
Attention, site généré par des moyens informatiques à partir de corpus bruts.
Les informations ne sont donc pas validées.

Embedded Formulas Extraction

Identifieur interne : 001C29 ( Main/Exploration ); précédent : 001C28; suivant : 001C30

Embedded Formulas Extraction

Auteurs : Afef Kacem [Tunisie] ; Abdel Belaïd [France] ; Mohamed Ben Ahmed [Tunisie]

Source :

RBID : Hal:inria-00099142

Descripteurs français

Abstract

A new approach for separating mathematics from usual text is presented. Contrary to the existing methods, it is more oriented toward the segmentation than the recognition, isolating the formulas outside and inside the text lines. The objective is to delimit a part of text which could disturb the OCR application, not yet trained for formula recognition and restructuring. The method is based on an adaptive segmentation working at two levels 1) A primary labelling identifies the more characteristic symbols; 2) A secondary labelling extends the context of the symbols for delimiting the formula inside the text.Experiments done on some commonly seen mathematical documents, show that our proposed method can achieve quite satisfactory rate making mathematical formulas extraction more feasible for real-world applications. The average rate of primary labelling of mathematical symbols is about 95.3% and their secondary labelling can improve the rate about 4%. Thus, about 95% of formulas are well extracted from images of documents printed with high quality

Url:


Affiliations:


Links toward previous steps (curation, corpus...)


Le document en format XML

<record>
<TEI>
<teiHeader>
<fileDesc>
<titleStmt>
<title xml:lang="en">Embedded Formulas Extraction</title>
<author>
<name sortKey="Kacem, Afef" sort="Kacem, Afef" uniqKey="Kacem A" first="Afef" last="Kacem">Afef Kacem</name>
<affiliation wicri:level="1">
<hal:affiliation type="laboratory" xml:id="struct-421707" status="VALID">
<orgName>Laboratoire RIADI-GDL [Manouba]</orgName>
<desc>
<address>
<addrLine>École Nationale des Sciences de l'Informatique (ENSI)Campus Universitaire de la Manouba,2010 Manouba, Tunisia</addrLine>
<country key="TN"></country>
</address>
<ref type="url">http://www.riadi.rnu.tn/</ref>
</desc>
<listRelation>
<relation active="#struct-301618" type="direct"></relation>
</listRelation>
<tutelles>
<tutelle active="#struct-301618" type="direct">
<org type="institution" xml:id="struct-301618" status="INCOMING">
<orgName>Ecole Nationale des Sciences de l'Informatique, Manouba, Tunisie</orgName>
<desc>
<address>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>Tunisie</country>
</affiliation>
</author>
<author>
<name sortKey="Belaid, Abdel" sort="Belaid, Abdel" uniqKey="Belaid A" first="Abdel" last="Belaïd">Abdel Belaïd</name>
<affiliation wicri:level="1">
<hal:affiliation type="researchteam" xml:id="struct-2362" status="OLD">
<orgName>READ</orgName>
<orgName type="acronym">READ</orgName>
<desc>
<address>
<country key="FR"></country>
</address>
</desc>
<listRelation>
<relation active="#struct-160" type="direct"></relation>
<relation name="UMR7503" active="#struct-441569" type="indirect"></relation>
<relation active="#struct-300009" type="indirect"></relation>
<relation active="#struct-300291" type="indirect"></relation>
<relation active="#struct-300292" type="indirect"></relation>
<relation active="#struct-300293" type="indirect"></relation>
</listRelation>
<tutelles>
<tutelle active="#struct-160" type="direct">
<org type="laboratory" xml:id="struct-160" status="OLD">
<orgName>Laboratoire Lorrain de Recherche en Informatique et ses Applications</orgName>
<orgName type="acronym">LORIA</orgName>
<desc>
<address>
<addrLine>Campus Scientifique BP 239 54506 Vandoeuvre-lès-Nancy Cedex</addrLine>
<country key="FR"></country>
</address>
<ref type="url">http://www.loria.fr</ref>
</desc>
<listRelation>
<relation name="UMR7503" active="#struct-441569" type="direct"></relation>
<relation active="#struct-300009" type="direct"></relation>
<relation active="#struct-300291" type="direct"></relation>
<relation active="#struct-300292" type="direct"></relation>
<relation active="#struct-300293" type="direct"></relation>
</listRelation>
</org>
</tutelle>
<tutelle name="UMR7503" active="#struct-441569" type="indirect">
<org type="institution" xml:id="struct-441569" status="VALID">
<idno type="IdRef">02636817X</idno>
<idno type="ISNI">0000000122597504</idno>
<orgName>Centre National de la Recherche Scientifique</orgName>
<orgName type="acronym">CNRS</orgName>
<date type="start">1939-10-19</date>
<desc>
<address>
<country key="FR"></country>
</address>
<ref type="url">http://www.cnrs.fr/</ref>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300009" type="indirect">
<org type="institution" xml:id="struct-300009" status="VALID">
<orgName>Institut National de Recherche en Informatique et en Automatique</orgName>
<orgName type="acronym">Inria</orgName>
<desc>
<address>
<addrLine>Domaine de VoluceauRocquencourt - BP 10578153 Le Chesnay Cedex</addrLine>
<country key="FR"></country>
</address>
<ref type="url">http://www.inria.fr/en/</ref>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300291" type="indirect">
<org type="institution" xml:id="struct-300291" status="OLD">
<orgName>Université Henri Poincaré - Nancy 1</orgName>
<orgName type="acronym">UHP</orgName>
<date type="end">2011-12-31</date>
<desc>
<address>
<addrLine>24-30 rue Lionnois, BP 60120, 54 003 NANCY cedex, France</addrLine>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300292" type="indirect">
<org type="institution" xml:id="struct-300292" status="OLD">
<orgName>Université Nancy 2</orgName>
<date type="end">2011-12-31</date>
<desc>
<address>
<addrLine>91 avenue de la Libération, BP 454, 54001 Nancy cedex</addrLine>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300293" type="indirect">
<org type="institution" xml:id="struct-300293" status="OLD">
<orgName>Institut National Polytechnique de Lorraine</orgName>
<orgName type="acronym">INPL</orgName>
<date type="end">2011-12-31</date>
<desc>
<address>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>France</country>
<placeName>
<settlement type="city">Nancy</settlement>
<region type="region" nuts="2">Lorraine</region>
</placeName>
<orgName type="university">Université Nancy 2</orgName>
<orgName type="institution" wicri:auto="newGroup">Université de Lorraine</orgName>
<placeName>
<settlement type="city">Nancy</settlement>
<region type="region" nuts="2">Lorraine</region>
</placeName>
<orgName type="university">Institut national polytechnique de Lorraine</orgName>
<orgName type="institution" wicri:auto="newGroup">Université de Lorraine</orgName>
</affiliation>
</author>
<author>
<name sortKey="Ben Ahmed, Mohamed" sort="Ben Ahmed, Mohamed" uniqKey="Ben Ahmed M" first="Mohamed" last="Ben Ahmed">Mohamed Ben Ahmed</name>
<affiliation wicri:level="1">
<hal:affiliation type="laboratory" xml:id="struct-421707" status="VALID">
<orgName>Laboratoire RIADI-GDL [Manouba]</orgName>
<desc>
<address>
<addrLine>École Nationale des Sciences de l'Informatique (ENSI)Campus Universitaire de la Manouba,2010 Manouba, Tunisia</addrLine>
<country key="TN"></country>
</address>
<ref type="url">http://www.riadi.rnu.tn/</ref>
</desc>
<listRelation>
<relation active="#struct-301618" type="direct"></relation>
</listRelation>
<tutelles>
<tutelle active="#struct-301618" type="direct">
<org type="institution" xml:id="struct-301618" status="INCOMING">
<orgName>Ecole Nationale des Sciences de l'Informatique, Manouba, Tunisie</orgName>
<desc>
<address>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>Tunisie</country>
</affiliation>
</author>
</titleStmt>
<publicationStmt>
<idno type="wicri:source">HAL</idno>
<idno type="RBID">Hal:inria-00099142</idno>
<idno type="halId">inria-00099142</idno>
<idno type="halUri">https://hal.inria.fr/inria-00099142</idno>
<idno type="url">https://hal.inria.fr/inria-00099142</idno>
<date when="2000-09">2000-09</date>
<idno type="wicri:Area/Hal/Corpus">000046</idno>
<idno type="wicri:Area/Hal/Curation">000046</idno>
<idno type="wicri:Area/Hal/Checkpoint">000153</idno>
<idno type="wicri:Area/Main/Merge">001D28</idno>
<idno type="wicri:Area/Main/Curation">001C29</idno>
<idno type="wicri:Area/Main/Exploration">001C29</idno>
</publicationStmt>
<sourceDesc>
<biblStruct>
<analytic>
<title xml:lang="en">Embedded Formulas Extraction</title>
<author>
<name sortKey="Kacem, Afef" sort="Kacem, Afef" uniqKey="Kacem A" first="Afef" last="Kacem">Afef Kacem</name>
<affiliation wicri:level="1">
<hal:affiliation type="laboratory" xml:id="struct-421707" status="VALID">
<orgName>Laboratoire RIADI-GDL [Manouba]</orgName>
<desc>
<address>
<addrLine>École Nationale des Sciences de l'Informatique (ENSI)Campus Universitaire de la Manouba,2010 Manouba, Tunisia</addrLine>
<country key="TN"></country>
</address>
<ref type="url">http://www.riadi.rnu.tn/</ref>
</desc>
<listRelation>
<relation active="#struct-301618" type="direct"></relation>
</listRelation>
<tutelles>
<tutelle active="#struct-301618" type="direct">
<org type="institution" xml:id="struct-301618" status="INCOMING">
<orgName>Ecole Nationale des Sciences de l'Informatique, Manouba, Tunisie</orgName>
<desc>
<address>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>Tunisie</country>
</affiliation>
</author>
<author>
<name sortKey="Belaid, Abdel" sort="Belaid, Abdel" uniqKey="Belaid A" first="Abdel" last="Belaïd">Abdel Belaïd</name>
<affiliation wicri:level="1">
<hal:affiliation type="researchteam" xml:id="struct-2362" status="OLD">
<orgName>READ</orgName>
<orgName type="acronym">READ</orgName>
<desc>
<address>
<country key="FR"></country>
</address>
</desc>
<listRelation>
<relation active="#struct-160" type="direct"></relation>
<relation name="UMR7503" active="#struct-441569" type="indirect"></relation>
<relation active="#struct-300009" type="indirect"></relation>
<relation active="#struct-300291" type="indirect"></relation>
<relation active="#struct-300292" type="indirect"></relation>
<relation active="#struct-300293" type="indirect"></relation>
</listRelation>
<tutelles>
<tutelle active="#struct-160" type="direct">
<org type="laboratory" xml:id="struct-160" status="OLD">
<orgName>Laboratoire Lorrain de Recherche en Informatique et ses Applications</orgName>
<orgName type="acronym">LORIA</orgName>
<desc>
<address>
<addrLine>Campus Scientifique BP 239 54506 Vandoeuvre-lès-Nancy Cedex</addrLine>
<country key="FR"></country>
</address>
<ref type="url">http://www.loria.fr</ref>
</desc>
<listRelation>
<relation name="UMR7503" active="#struct-441569" type="direct"></relation>
<relation active="#struct-300009" type="direct"></relation>
<relation active="#struct-300291" type="direct"></relation>
<relation active="#struct-300292" type="direct"></relation>
<relation active="#struct-300293" type="direct"></relation>
</listRelation>
</org>
</tutelle>
<tutelle name="UMR7503" active="#struct-441569" type="indirect">
<org type="institution" xml:id="struct-441569" status="VALID">
<idno type="IdRef">02636817X</idno>
<idno type="ISNI">0000000122597504</idno>
<orgName>Centre National de la Recherche Scientifique</orgName>
<orgName type="acronym">CNRS</orgName>
<date type="start">1939-10-19</date>
<desc>
<address>
<country key="FR"></country>
</address>
<ref type="url">http://www.cnrs.fr/</ref>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300009" type="indirect">
<org type="institution" xml:id="struct-300009" status="VALID">
<orgName>Institut National de Recherche en Informatique et en Automatique</orgName>
<orgName type="acronym">Inria</orgName>
<desc>
<address>
<addrLine>Domaine de VoluceauRocquencourt - BP 10578153 Le Chesnay Cedex</addrLine>
<country key="FR"></country>
</address>
<ref type="url">http://www.inria.fr/en/</ref>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300291" type="indirect">
<org type="institution" xml:id="struct-300291" status="OLD">
<orgName>Université Henri Poincaré - Nancy 1</orgName>
<orgName type="acronym">UHP</orgName>
<date type="end">2011-12-31</date>
<desc>
<address>
<addrLine>24-30 rue Lionnois, BP 60120, 54 003 NANCY cedex, France</addrLine>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300292" type="indirect">
<org type="institution" xml:id="struct-300292" status="OLD">
<orgName>Université Nancy 2</orgName>
<date type="end">2011-12-31</date>
<desc>
<address>
<addrLine>91 avenue de la Libération, BP 454, 54001 Nancy cedex</addrLine>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
<tutelle active="#struct-300293" type="indirect">
<org type="institution" xml:id="struct-300293" status="OLD">
<orgName>Institut National Polytechnique de Lorraine</orgName>
<orgName type="acronym">INPL</orgName>
<date type="end">2011-12-31</date>
<desc>
<address>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>France</country>
<placeName>
<settlement type="city">Nancy</settlement>
<region type="region" nuts="2">Lorraine</region>
</placeName>
<orgName type="university">Université Nancy 2</orgName>
<orgName type="institution" wicri:auto="newGroup">Université de Lorraine</orgName>
<placeName>
<settlement type="city">Nancy</settlement>
<region type="region" nuts="2">Lorraine</region>
</placeName>
<orgName type="university">Institut national polytechnique de Lorraine</orgName>
<orgName type="institution" wicri:auto="newGroup">Université de Lorraine</orgName>
</affiliation>
</author>
<author>
<name sortKey="Ben Ahmed, Mohamed" sort="Ben Ahmed, Mohamed" uniqKey="Ben Ahmed M" first="Mohamed" last="Ben Ahmed">Mohamed Ben Ahmed</name>
<affiliation wicri:level="1">
<hal:affiliation type="laboratory" xml:id="struct-421707" status="VALID">
<orgName>Laboratoire RIADI-GDL [Manouba]</orgName>
<desc>
<address>
<addrLine>École Nationale des Sciences de l'Informatique (ENSI)Campus Universitaire de la Manouba,2010 Manouba, Tunisia</addrLine>
<country key="TN"></country>
</address>
<ref type="url">http://www.riadi.rnu.tn/</ref>
</desc>
<listRelation>
<relation active="#struct-301618" type="direct"></relation>
</listRelation>
<tutelles>
<tutelle active="#struct-301618" type="direct">
<org type="institution" xml:id="struct-301618" status="INCOMING">
<orgName>Ecole Nationale des Sciences de l'Informatique, Manouba, Tunisie</orgName>
<desc>
<address>
<country key="FR"></country>
</address>
</desc>
</org>
</tutelle>
</tutelles>
</hal:affiliation>
<country>Tunisie</country>
</affiliation>
</author>
</analytic>
</biblStruct>
</sourceDesc>
</fileDesc>
<profileDesc>
<textClass>
<keywords scheme="mix" xml:lang="fr">
<term>fuzzy logic mathematics segmentation</term>
<term>logique floue</term>
<term>segmentation de documents mathématiques</term>
</keywords>
</textClass>
</profileDesc>
</teiHeader>
<front>
<div type="abstract" xml:lang="en">A new approach for separating mathematics from usual text is presented. Contrary to the existing methods, it is more oriented toward the segmentation than the recognition, isolating the formulas outside and inside the text lines. The objective is to delimit a part of text which could disturb the OCR application, not yet trained for formula recognition and restructuring. The method is based on an adaptive segmentation working at two levels 1) A primary labelling identifies the more characteristic symbols; 2) A secondary labelling extends the context of the symbols for delimiting the formula inside the text.Experiments done on some commonly seen mathematical documents, show that our proposed method can achieve quite satisfactory rate making mathematical formulas extraction more feasible for real-world applications. The average rate of primary labelling of mathematical symbols is about 95.3% and their secondary labelling can improve the rate about 4%. Thus, about 95% of formulas are well extracted from images of documents printed with high quality</div>
</front>
</TEI>
<affiliations>
<list>
<country>
<li>France</li>
<li>Tunisie</li>
</country>
<region>
<li>Lorraine</li>
</region>
<settlement>
<li>Nancy</li>
</settlement>
<orgName>
<li>Institut national polytechnique de Lorraine</li>
<li>Université Nancy 2</li>
<li>Université de Lorraine</li>
</orgName>
</list>
<tree>
<country name="Tunisie">
<noRegion>
<name sortKey="Kacem, Afef" sort="Kacem, Afef" uniqKey="Kacem A" first="Afef" last="Kacem">Afef Kacem</name>
</noRegion>
<name sortKey="Ben Ahmed, Mohamed" sort="Ben Ahmed, Mohamed" uniqKey="Ben Ahmed M" first="Mohamed" last="Ben Ahmed">Mohamed Ben Ahmed</name>
</country>
<country name="France">
<region name="Lorraine">
<name sortKey="Belaid, Abdel" sort="Belaid, Abdel" uniqKey="Belaid A" first="Abdel" last="Belaïd">Abdel Belaïd</name>
</region>
</country>
</tree>
</affiliations>
</record>

Pour manipuler ce document sous Unix (Dilib)

EXPLOR_STEP=$WICRI_ROOT/Ticri/CIDE/explor/OcrV1/Data/Main/Exploration
HfdSelect -h $EXPLOR_STEP/biblio.hfd -nk 001C29 | SxmlIndent | more

Ou

HfdSelect -h $EXPLOR_AREA/Data/Main/Exploration/biblio.hfd -nk 001C29 | SxmlIndent | more

Pour mettre un lien sur cette page dans le réseau Wicri

{{Explor lien
   |wiki=    Ticri/CIDE
   |area=    OcrV1
   |flux=    Main
   |étape=   Exploration
   |type=    RBID
   |clé=     Hal:inria-00099142
   |texte=   Embedded Formulas Extraction
}}

Wicri

This area was generated with Dilib version V0.6.32.
Data generation: Sat Nov 11 16:53:45 2017. Site generation: Mon Mar 11 23:15:16 2024